Papers with cross-modal pre-training

2 papers
Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to understand chart plots are difficult to apply to visual-language tasks.
Approach: They propose a V+L model that learns how to interpret table information from chart images via cross-modal pre-training on plot table pairs.
Outcome: The proposed model outperforms state-of-the-art models on the chartQA benchmark by over 8% performance gains.
MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for Large Multimodal Models (LMMs) are constrained by static representations, inadequately evaluating their ability to understand time-sensitive knowledge.
Approach: They propose a benchmark containing 2,104 time-sensitive knowledge samples spanning six knowledge types to evaluate temporal awareness along 6 key dimensions and 11 challenging tasks.
Outcome: The proposed benchmark measures temporal awareness along 6 key dimensions and 11 tasks, while most open-source LMMs still lack time understanding ability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations